Papers with statistical significance testing

7 papers
NLPStatTest: A Toolkit for Comparing NLP System Performance (2020.aacl-demo)

Copied to clipboard

Challenge: Statistical significance testing is used to compare NLP system performance, but p-values alone are insufficient because statistical significance differs from practical significance.
Approach: They propose a three-stage procedure for comparing NLP system performance and a toolkit that automates the process.
Outcome: The proposed procedure is based on a three-stage procedure and compares it with existing statistical testing toolkits.
A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media (2024.lrec-main)

Copied to clipboard

Challenge: Dehumanisation involves the perception and/or treatment of a social group’s members as less than human.
Approach: They propose to use a new sentiment resource to make it easier to transfer to other languages and to evaluate and use . they then apply the method to study attitudes to migration expressed in Slovene newspapers, and examine how this discourse changed between the 2015-16 migration crisis and the 2022-23 period following the war in Ukraine.
Outcome: The proposed method is easier to transfer to other languages and evaluates . it combines zero-shot cross-lingual valence and arousal detection with statistical significance testing to examine attitudes to migration expressed in Slovene newspapers .
Methods, Applications, and Directions of Learning-to-Rank in NLP Research (2024.findings-naacl)

Copied to clipboard

Challenge: Learning-to-rank (LTR) algorithms aim to order items according to some criteria.
Approach: They focus on the formal background of LTR and the most widely-used supervised methods . they also discuss how large language models are changing the LTR landscape .
Outcome: The proposed methods are used in natural language processing and information retrieval tasks.
The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing (P18-1)

Copied to clipboard

Challenge: Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.
Approach: They propose a protocol for statistical significance test selection in NLP setups . they propose he proposes a survey of the most relevant tests to help guide the protocol .
Outcome: The proposed protocol includes a survey of the most relevant tests.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems (2022.acl-long)

Copied to clipboard

Challenge: Evaluation of open-domain dialogue systems is challenging and unreliable . human evaluation of live conversations is highly reliable, but reliability cannot be assumed .
Approach: They propose a method of open-domain dialogue evaluation that is highly reliable . they compare live conversations with models that avoid pre-created reference dialogues .
Outcome: The proposed method is highly reliable while remaining feasible and low cost.
Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers (2021.acl-long)

Copied to clipboard

Challenge: a meta-evaluation of machine translation (MT) has been conducted in 769 research papers . a recent study shows that evaluation practices have changed over the past decade .
Approach: They propose a meta-evaluation method for machine translation that uses BLEU scores to evaluate MT performance.
Outcome: The proposed meta-evaluation of machine translation shows that evaluation practices have changed over the past decade . the authors suggest that the evaluation process should be streamlined and standardized to ensure the validity of the evaluation method .
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)

Copied to clipboard

Challenge: comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority .
Approach: They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance.
Outcome: The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations